Papers with multilingual robustness
VLURes: Benchmarking Long-Text Grounding and Cross-Lingual Robustness in Vision Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | ***VLURes** provides a practical testbed for long-text grounding and multilingual robustness in web-realistic agent settings. |
| Approach: | They propose a multilingual benchmark for evaluating vision-language models under long-text grounding. |
| Outcome: | ***VLURes** provides a testbed for long-text grounding and multilingual robustness in web-realistic agent settings. |
PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts (2026.acl-long)
Copied to clipboard
| Challenge: | Large Reasoning Models (LRMs) are embedded in agentic frameworks and are under-evaluated. |
| Approach: | They propose a multilingual benchmark for agentic information synthesis using PolitNuggets . they standardize evaluation with an optimized Supervisor–Searcher multi-agent system . |
| Outcome: | The proposed model can discover and synthesize "long-tail" facts from dispersed sources. |